Papers with distance measures
ANALOGICAL - A Novel Benchmark for Long Text Analogy Evaluation in Large Language Models (2023.findings-acl)
Copied to clipboard
Thilini Wijesiriwardene, Ruwan Wickramarachchi, Bimal Gajera, Shreeyash Gowaikar, Chandan Gupta, Aman Chadha, Aishwarya Naresh Reganti, Amit Sheth, Amitava Das
| Challenge: | Modern large language models are evaluated on extrinsic measures based on benchmarks such as GLUE and SuperGLUE. |
| Approach: | They propose a benchmark to intrinsically evaluate large language models across a taxonomy of analogies of long text with six levels of complexity. |
| Outcome: | The proposed benchmark evaluates LLMs across a taxonomy of analogies of long text with six levels of complexity. |
Computing with Subjectivity Lexicons (2020.lrec-1)
Copied to clipboard
Caio L. M. Jeronimo, Claudio E. C. Campelo, Leandro Balby Marinho, Allan Sales, Adriano Veloso, Roberta Viola
| Challenge: | a new set of lexicons for expressing subjectivity in text documents is presented . lexiconics are useful resources for identifying semantics relevant to sentiment, emotion, personality, language bias, mood, and attitude. |
| Approach: | They propose a set of lexicons for expressing subjectivity in Brazilian Portuguese text documents . they use word embedding techniques to capture semantically related words to the ones in the lexicos . |
| Outcome: | The proposed lexicons represent different subjectivity dimensions and are more compact in number of terms. |
Prediction Hubs are Context-Informed Frequent Tokens in LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | Hubness is a tendency for a few points to be among the nearest neighbours of a disproportionate number of other points. |
| Approach: | They show that only large-scale representation comparisons are not characterized by hubness . they show that hubs are the result of context-modulated frequent tokens . |
| Outcome: | The results show that the comparison between context and unembedding vectors does not result in hubness . the findings suggest that hubness is not a negative property that needs to be mitigated when LLMs are being used for next token prediction. |